Proteins: Structure, Function, and Bioinformatics
○ Wiley
Preprints posted in the last 30 days, ranked by how well they match Proteins: Structure, Function, and Bioinformatics's content profile, based on 88 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.
Eicholt, L. A.; Middendorf, L.
Show abstract
Structure and disorder predictors are increasingly used as decision-grade tools in protein engineering and in the analysis of newly emerged proteins, yet how the current state-of-the-art behaves on sequences outside the well-charted evolutionary space remains poorly characterised. We previously reported that AlphaFold2 confidence and the disorder predictor flDPnn produced discordant predictions for naturally evolved de novo Drosophila proteins and for shuffled sequences. Here, we revisit the comparison with AlphaFold3 and the best-performing disorder predictor PUNCH2 on the same sequence sets together with conserved Drosophila proteins and intrinsically disordered proteins. The discordance persists: pLDDT correlates positively with PUNCH2 disorder in random and de novo proteins and negatively with {beta}-strand fraction, opposite to the conserved and disordered baselines. A class-specific, score-defined driver subset jointly captures the unusual high-pLDDT, high-disorder, low-strand combination and contains 24.5% of de novo, 29.4% of random, 5.1% of conserved, and 1.3% of disordered proteins. Removing this subset normalises the correlations. A held-out classifier trained on architectural and compositional features that were not used in the driver definition recovers the subset, with helix and coil fraction, sequence length, entropy and hydropathy as the strongest predictors. The discordance is therefore not a sequence-class artefact but a localised, compositionally identifiable phenotype that current predictors handle in a non-canonical way - a concrete failure mode that protein designers and others working on sequences remote in sequence space should be aware of when relying on predictor outputs.
Kudo, T.; Ekimoto, T.; Yamane, T.; Ikeguchi, M.
Show abstract
Many functional RNA motifs adopt structures that deviate from the canonical A-form helix and are emerging targets for RNA-directed therapeutics. The microtubule-associated protein tau (MAPT) A-bulge motif (5'-GCAGU/5'-ACGU) is one such motif. Because its structure is stabilized by a delicate balance of local interactions, its accurate modeling remains a major challenge for molecular dynamics (MD) simulations. The experimentally determined nuclear magnetic resonance (NMR) structure of the MAPT A-bulge motif provides a stringent test of whether RNA force fields can accurately reproduce the experimentally observed conformation. Most current AMBER-family RNA force-field models have incorrectly favored a non-native base-triple state of the MAPT A-bulge motif over the experimentally observed stacked state. Structural comparison of the stacked and base-triple conformations revealed that overly favorable NH-N hydrogen bonds between the bulged adenosine and an adjacent Watson-Crick base pair were the primary source of this imbalance. We developed gHBfix-18Ab, an 18-component hydrogen-bond correction that distinguishes NH and NH2; donors. gHBfix-18Ab was combined with the previously developed OL3CP and NBfix0BPh corrections to generate the composite model gHBfix-18Ab*. This model restored the experimentally observed stacked state as the global minimum in the calculated free-energy profile and improved agreement with NMR-derived distance data for the A-bulge region. Importantly, gHBfix-18Ab* did not produce marked structural destabilization of the cUUCGg tetraloop, a widely used benchmark for RNA force-field validation, suggesting that the refinement preserves the stability of the unrelated RNA motif. These results demonstrate that targeted refinement of hydrogen-bond interactions provides a practical strategy for systematic improvement of RNA force fields toward more accurate modeling of noncanonical RNA motifs.
Sartori, J.; Guimaraes, A. C. R.; Machado, L. d. A.
Show abstract
Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap, we introduce EnzymARC, a novel benchmark dataset of putative non-functional decoy sequences generated via structure-guided, systematic disruption of active sites (targeting catalytic residues and surrounding 5 A, 10 A, and 15 A radii) from experimentally annotated enzymes. We evaluated three distinct prediction paradigms against this dataset: homology-based annotation (DIAMOND), contrastive learning with protein language models (CLEAN), and a deep learning model incorporating non-enzyme discrimination (DeepEC). Our findings reveal that current models are highly vulnerable to phylogenetic shortcuts. Both DIAMOND and CLEAN exhibited false positive rates exceeding 90\% for low-perturbation decoys, confidently assigning the original EC numbers despite the destruction of the catalytic machinery. While DeepEC demonstrated improved sensitivity at higher perturbation levels, highlighting the benefit of negative training examples, all models struggled to identify targeted active-site disruptions. We demonstrate that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and we propose that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.
Miyaguchi, I.; Hata, H.; Kuribayashi, T.; Takahashi, S.; Kashima, A.; Murasaki, K.; Matsumoto, S.; Terayama, K.; Ohta, M.; Ikeguchi, M.
Show abstract
Accurate assessment of ligand coordinate-density consistency across different resolutions remains challenging in macromolecular crystallography. We introduce the atomic Box Correlation Coefficient (aBCC), an atom-level metric for evaluating the consistency between ligand atomic coordinates and electron density in a resolution-standardized framework. To predict aBCC values from electron-density maps, we developed QAEmap, a machine-learning model based on three-dimensional convolutional neural networks (3D-CNNs). The model was trained using Fourier-truncated electron-density maps and corresponding ligand coordinates generated from high-resolution structures in the Protein Data Bank. It was evaluated using both Fourier-truncated electron-density maps and experimentally determined PDB structures. was evaluated using both Fourier-truncated electron-density maps and experimentally determined PDB structures.The prediction accuracy gradually decreased with decreasing resolution, but remained reliable up to [~]3.5 [A]. These results demonstrate that aBCC enables resolution-standardized atom-wise evaluation of coordinate-density consistency across different resolutions and provide a foundation for further development and refinement of machine learning-based coordinate validation. SynopsisWe introduce the atomic box correlation coefficient (aBCC), a machine learning-based metric for the resolution-standardized atom-level evaluation of ligand coordinate-density consistency in crystallographic structures. aBCC provides a common framework for assessing and communicating the local coordinate reliability between structural biologists and researchers in structure-based drug discovery.
Refaee, A. A.; Milanetti, E.; Roeder, K.; Ruocco, G.; Iacoangeli, A.
Show abstract
Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease characterised by progressive motor neuron degeneration. Mutations in the SOD1 gene represent the second most common genetic cause of ALS (ALS), and distinct SOD1 missense variants present with markedly different clinical profiles. A4V leads to an aggressive form of the disease (median survival [~]1y), H46R confers a mild, slowly progressive course and I113T exhibits an intermediate phenotype. The molecular basis by which these mutations produce divergent clinical outcomes remains poorly understood. We performed extensive classical molecular dynamics simulations of wild-type SOD1 and the three ALS-associated variants in the apo monomeric state to attempt to investigate the mechanisms behind such phenotypic differences. Structural stability, global compactness, and conformational flexibility, as well as analysis of collective motions between residues and estimation of free energy, were assessed. The H46R, A4V, and I113T variants exhibited distinct dynamic behaviours, highlighting differences in structural stability, local flexibility, and intramolecular interactions. These findings suggest that specific structural regions may contribute differently to protein dysfunction and could represent key elements for understanding the relationship between molecular dynamic properties and the differing clinical severity associated with these variants. Most strikingly, H46R exhibited exceptional structural stability across every analytical level, the lowest global deviation, most attenuated local flexibility, strongest internal dynamic coordination, and the deepest, most confined free energy basins of any system examined. This convergent multi-layered evidence of structural restraint provides a compelling mechanistic basis for the mild and slowly progressive clinical course of H46R ALS, suggesting that enhanced conformational rigidity, rather than bulk destabilisation, is the defining biophysical feature of this variant, and that its pathogenic mechanism operates through a route fundamentally decoupled from the aggregation-driven toxicity that characterises the more aggressive SOD1-ALS mutations.
Takahashi, N.; Abe, N.; Mabuchi, T.; Fukuyama, M.; Terauchi, Y.; Tanaka, T.; Yoshimi, A.; Yabu, H.; Abe, K.
Show abstract
Hydrophobins are biosurfactant proteins that coat the cell surfaces of filamentous fungi. On the conidial surface, hydrophobins self-assemble into rodlets, forming a dense hydrophobic film that promotes air-dispersibility. Although rodlet formation is closely associated with the physiology of filamentous fungi, its underlying molecular mechanisms remain largely unknown. Previously, we revealed that RolA, a hydrophobin derived from Aspergillus oryzae, forms rodlets at the air-water interface. In this study, we focused on the flexible N-terminal region of RolA, which lacks a well-defined tertiary structure, and hypothesized that this intrinsically disordered region regulates rodlet formation. To investigate its role, we used RolA mutants with reduced charges in the N-terminal region and analyzed the rodlet formation process on the surface of a water-in-air sessile droplet using atomic force microscopy. In addition, we quantitatively characterized rodlet formation at the air-water interface by applying a kinetic perspective to the interfacial tension change profiles obtained from dynamic surface tension measurements. The results suggested that RolA first forms a monolayer at the air-water interface, then rodlet formation proceeds through the continuous supply of free RolA monomers from the bulk phase to the interfacial RolA film. Our molecular dynamics simulations of RolA at the interface supported a model in which RolA molecules within the interfacial film interact with free monomers in the bulk phase through their N-terminal regions. These results reveal a previously unidentified role of the N-terminal region in rodlet formation and provide a more comprehensive framework for understanding the molecular mechanism underlying RolA rodlet formation.
Simpkin, A. J.; Johnson, E.; Rigden, D.
Show abstract
Motivation: The actual interface pTM score (actifpTM) is a modified version of the ipTM score that limits the calculation to only those residues at the interface. Whilst actifpTM provides an effective interface quality score, a limiting factor is that it makes use of the predicted aligned error (PAE) with probabilities, information that is generated during a ColabFold run, but not output by the package or other model prediction software. The consequent inability to generate actifpTM scores for the results of software such as AlphaFold 2 or AlphaFold 3 has limited its adoption. With reactifpTM we address this problem by providing a standalone tool that can be run on the standard outputs of most model prediction packages. Results: Using the same underlying principles as actifpTM, reactifpTM has been developed to use standard output files from model prediction software (a model and corresponding PAE) to perform an actifpTM-like calculation. ColabFold models were generated for a dataset of 1079 known interfaces in the PDB. A strong correlation was shown between actifpTM and reactifpTM for this dataset. Availability and implementation: reactifpTM is coded in Python. All scripts and associated documentation are available from https://github.com/hlasimpk/reactifptm or https://pypi.org/project/reactifptm.
Bromley, A. C.; Kruse, N. A.; Brower, C. R.; Beam, M. K.; Hammer, N. I.; Fortenberry, R. C.; Reinemann, D. N.
Show abstract
This present work shows that E-hook fragments possess functional structure differences governed by electrostatic interactions and sequence composition. The acidic C-terminal tails of tubulin, known as E-hooks, play a central role in regulating interactions between microtubules and motor proteins, microtubule-associated proteins, and enzymatic modifiers. Despite their functional importance, the intrinsic structural properties of these peptide segments remain poorly characterized due to their intrinsically disordered nature. In this work, we present quantum-mechanically optimized structures of hexamer peptides derived from {beta}-tubulin E-hook sequences. Density functional theory calculations were used to optimize peptide geometries using progressively larger basis sets. From the optimized geometries we calculated theoretical Raman spectra, Ramachandran backbone dihedral distributions, and measured radii of gyration to resolve composition dependent structural tendencies. The combined Raman and conformational analyses provide a systematic computational approach for comparing simulated and experimental Raman spectra of tubulin E-hooks and other intrinsically disordered proteins and offer insight into how E-hooks contribute to the recognition mechanisms underlying the tubulin code.
LARUE, V.; Nonin-Lecomte, S.
Show abstract
We present the solution structures of HIV-1 proteins NC(p7)1-55 corresponding to the full-length NC(p7) and mature p6. The studies were carried in water and, to mimic the membrane, in micellar DPC (Dodecylphosphocholine) conditions. Our results unravel for the first time the structure adopted by the N-terminal amino acids of the free NC(p7)1-55, with the formation of a small helix spanning residues F6 to R10. Our NMR and Fluorescence Anisotropy data disclose an interaction between NC(p7)1-55 and p6 both in water and DPC, with respective Kd of 2.5mM and 370 mM at 23{degrees}C. The interaction is thus strengthened in lipidic conditions. Protein p6 stabilizes the N-terminus of NC(p7)1-55 while increasing at the same time the dynamic of the first zinc finger. Although the entire p6 sequence is involved in the interaction, we show that its C-terminal region is particularly sensitive to the presence of NC(p7)1-55, with a propensity of forming a a helix ranging from amino acids S111 to F116. This study brings experimental evidence of a direct protein-protein interaction between p6 and the N-terminal region of NC(p7)1-55. We further show that such interaction is readily accommodated within the NC(p15) framework and hypothesize that it may facilitate the selective assembly of assembly of the viral genomic RNA (gRNA) in the cell.
Riepenhausen, L.; Costa, F.; Andreeva, A.; Bateman, A.
Show abstract
Motivation: Continuing advances in genome and metagenome sequencing expand the number of identified conserved protein families that remain functionally uncharacterized and contain domains of unknown function (DUFs). Functional-association resources such as STRING provide biological context, but mostly do not distinguish indirect association from physical interaction. We assessed whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins. Results: We generated four structural-prediction cohorts from STRING associations involving DUF-containing proteins and evaluated the predicted complexes using interface ipSAE, average pLDDT and buried surface area. An L2-regularized logistic regression model was trained on an initial cohort of predictions from high-confidence STRING associations to prioritize DUF-containing candidates likely to produce structurally confident AlphaFold 3 complexes. The model was then applied across all 12,535 organisms represented in STRING v12.0, followed by grouping into DUF-family and partner-architecture modules, covering 2,076 unique DUF families. The final L2-model screen contained 12,298 successfully modelled protein pairs, including 1,208 (9.82%) complexes meeting a strict-confidence criterion and 2,433 (19.78%) meeting a more liberal confidence criterion. Two examples suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation. Availability and implementation: Predicted structures and associated metadata are available through Zenodo at https://doi.org/10.5281/zenodo.21875362. The model implementation and code used to generate the analyses and figures are available at https://github.com/linoriep/Proteome-scale-structure-prediction-of-DUF-containing-protein-protein-interactions.
Li, Z.; Yuan, Y.; Hu, K.; Pan, P.; He, F.
Show abstract
Cyclic peptides are a rapidly expanding class of therapeutics, but the reliability of deep-learning structure prediction for cyclic peptide-protein complexes has not been systematically evaluated. We assembled a curated benchmark of 111 nonredundant complexes spanning five cyclization chemistries and assessed two co-folding models, Boltz and Protenix, each generating 100 poses per target (22,200 total). Stratifying all poses by complex attributes, we found that disulfidecyclized peptides and small protein targets (200 or fewer target residues) were predicted significantly worse by both tools, with target size the largest and most consistent effect; overall accuracy nevertheless remained high (median top-pose DockQ of about 0.89, 96-98% of targets Acceptable or better), indicating that pose generation is rarely the bottleneck. Conversely, native model ranking scores correlated only moderately with pose quality (Spearman rank correlations of 0.53-0.66): approximately 12% of poses showed high model ranking score/confidence despite poor pose DockQ quality, and the highest-quality pose was not ranked first for nearly every target. We therefore augmented the native score with externally computed interface descriptors normalized by chain length, principally the per-residue density of inter-chain hydrogen bonds, in a gradient-boosted rescoring model evaluated under target-grouped cross-validation that prevents leakage, improving out-of-fold ROC-AUC for both tools, significantly so for Protenix. Together, these findings identify pose ranking, rather than pose generation, as the major limitation of current cyclic peptide-protein complex prediction and demonstrate that complementary structural features can improve confidence-based pose selection.
Friedl, A.; Manst, D.
Show abstract
Background: Comparisons between independently predicted wild-type and missense-variant protein structures can generate mechanistic hypotheses, but small apparent differences may reflect model-selection variability rather than mutation-specific effects. Methods: Human mitochondrial DNA polymerase gamma (POLG; UniProt P54098) variants p.Arg627Gln (R627Q) and p.Trp748Ser (W748S) were evaluated using five AlphaFold2-PTM network-model outputs per condition generated with one random seed under matched ColabFold settings. Ten pairwise wild type comparisons at each site described between-network model-selection variability. Variant effects were summarized across five within-network wild-type-versus-variant comparisons using rotation-invariant local C-alpha pair distances and local displacement after global and local alignment. Because these comparison designs differ, the wild-type distribution was used as context rather than a mutation-effect null. Wild-type cryo-EM structure 9GGF was used for contact and interface mapping. Experimental A467T and G848S structures 9GGE and 9GGC provided contextual benchmarks. Results: R627Q measurements fell within the range of between-network wild-type differences: its median mean local pair-distance change was 0.170 angstrom, compared with a wild-type median of 0.170 angstrom, and its locally aligned displacement was 0.265 versus 0.248 angstrom. W748S showed higher median values (0.168 versus 0.132 angstrom for pair-distance change; 0.236 versus 0.182 angstrom for locally aligned displacement), but the ranges overlapped and the comparison-design asymmetry precluded a calibrated mutation-effect percentile. Experimental A467T and G848S comparisons produced local changes of similar magnitude. In 9GGF, R627 and W748 directly shared a local microenvironment, with a minimum heavy-atom distance of 3.53 angstrom. R627 also formed short polar-contact candidates with D629 and D743, whereas W748 occupied a hydrophobic packing environment containing Y622 and F750. Both sites were more than 18 angstrom from nucleic acid, more than 30 angstrom from POLG2, and more than 33 angstrom from PZL-A in a ligand-bound structure. Conclusions: Available AlphaFold2 comparisons do not establish a mutation-specific structural deformation for either variant. Experimental-structure mapping supports testable physicochemical hypotheses involving a shared R627-W748 microenvironment - loss of an arginine-centered polar network for R627Q and disruption of a buried aromatic environment for W748S - but not direct DNA, POLG2, or PZL-A contact mechanisms. Matched control substitutions and independent seeds are required to calibrate small mutation-associated structural deltas.
Liu, D.; Sreenivasan, S.; Gray, C. J.; Cleveland, H. C.; Swint-Kruse, L.
Show abstract
A central challenge in molecular biology is understanding how amino acid substitutions modulate various features of protein function and stability. To illuminate the complexities of this relationship, high-throughput (HTP) assays are increasingly used to assess site-saturating mutagenesis libraries. A common downstream analysis is to average the set of twenty outcomes at each amino acid position for comparison with structural and evolutionary features. Average values clearly identify positions that tolerate most substitutions (neutral positions) and positions where most substitutions abolish activity (toggle positions). However, average values conceal the existence of rheostat positions, where different amino acid substitutions sample a wide range of outcomes. To quantitatively identify rheostat positions, we previously developed a histogram-based analysis that we here expand by: (i) incorporating new position classes observed in experimental studies of rheostat positions; (ii) formalizing a hierarchy of class assignments; (iii) refining error-based identification of neutral positions; and (iv) statistically assessing the robustness of class assignments to changes in experimental and computational parameters. RheoScale 2.0 is implemented in Excel and newly implemented in Python for facile integration with existing HTP pipelines; all parameters are customizable. Example analyses are shown for three HTP datasets of the SARS-CoV-2 papain-like protease. Results illustrate two aspects that influence interpretation of HTP data: First, position assignments (and substitution outcomes) depend highly on the measured feature. Second, many protein positions play multiple roles in the sequence-structure-function relationship. The recognition of varied position roles will advance understanding of pathogen evolution, protein engineering, and variant interpretation for personalized medicine. SummaryRheoScale 2.0 improves how high-throughput mutational data are interpreted by identifying protein positions where amino acid substitutions act like biological dimmer switches. By enabling more nuanced assignment of position behavior, beyond neutral or deleterious outcomes, this analysis framework advances studies of sequence-structure-function relationships and has broad relevance for understanding protein evolution, engineering proteins with desired properties, and interpreting variants linked to human disease. SOFTWARE AVAILABILITYhttps://github.com/liskinsk/RheoScale-calculator
Moore, C. W.
Show abstract
Methods for predicting cryptic binding sites are compared almost exclusively on top-n recovery, a number that conflates two independent abilities: proposing a candidate at the right location, and ranking it highly enough to be seen. We separate them by retaining the per-candidate overlap of every proposal, rather than only the top five, for four structurally different detectors spanning 2009 to 2026, across the CryptoBench benchmark. The separation is large and it reorders the field. On the designated test fold of 178 structures, fpocket, a purely geometric method from 2009, proposes a qualifying candidate for 74.2% of targets, the highest coverage of any tool tested, yet surfaces one in its top five for only 43.8%. P2Rank proposes qualifying candidates for 66.3% and surfaces 63.5%, and IF-SitePred, a 2024 method built on protein language model embeddings, proposes 70.8% and surfaces 61.8%. Coverage across tools varies by 8 points while conversion, the share of a tools own coverage that reaches the top five, varies from 59% to 96%. Unioning the four detectors reaches 92.1% coverage, and only 7.9% of cryptic sites are invisible to all of them. The fields headroom is therefore predominantly in ranking and in combination, not in detection: perfect ranking of a single tools existing proposals would reach 74.2%, and of the union 92.1%, against the 66.3% currently achieved. We show the practical consequence is governed by candidate budget. Added coverage converts to recovery at about 85% while a structure carries fewer than roughly fifteen candidates and at about 51% above it, which explains a series of interventions that raised coverage and returned nothing. Working within that budget, proposing pockets from a protein language model at locations where geometry finds no concavity improves single-structure recovery by 8.5% (95% CI +4.0 to +13.6) on test-fold data, and lets a five-conformer ensemble match a twenty-conformer one at a third of the wall clock. We release per-candidate overlaps for all tools so that coverage and conversion can be reported separately without re-running any method.
Garimella, S. C.; Bhargava, Y.
Show abstract
Autosomal dominant neovascular inflammatory vitreoretinopathy (ADNIV) is a rare retinal disease caused by gain-of-function mutations in the non-classical calcium-activated cysteine protease calpain-5 (CAPN5). These mutations lower the calcium threshold for catalytic-triad alignment with downstream effects including excessive proteolysis and retinal degeneration, making CAPN5 a therapeutic target. Clinical studies showed that knockout of calpain-5 resulted in no negative side effects, supporting therapy through inhibition. We mapped the druggable pockets of CAPN5 with a 500 ns phenol cosolvent molecular dynamics (MD) simulation. Occupancy analysis resolved five pockets, against which 448,314 COCONUT natural products were screened with Uni-Dock (2,241,570 docked combinations). In parallel, BoltzGen was used to design peptide binders against multiple candidate regions, from which three were selected: the PC1-PC2 subdomain interface, the PC2 regulatory loop (PC2L1) and the catalytic region. The top three designs were co-folded with Boltz-2 at high interface confidence (ipTM 0.91-0.95). The top three peptides and four small molecules were then simulated against wild-type CAPN5 and the four canonical ADNIV variants R243L, L244P, K250N and R289W, each condition in independent triplicate, giving 105 production simulations of 100 ns. Scoring by MM-PBSA revealed favorable peptide interface energies, the most favorable being the largest of the three designs ({Delta}TOTAL -58.6 {+/-} 5.6 kcal/mol for a 23-residue peptide against wild type), while the small-molecule panel returned -8.7 to -23.8 kcal/mol. A total of 15.8 {micro}s of cosolvent, filtering, and production MD prioritizes the catalytic cleft and an adjacent groove for experimental testing and provides candidate peptide and small-molecule binders for evaluating CAPN5 inhibition in ADNIV.
Seker, A.; Anand, S.; Marintchev, A.
Show abstract
Eukaryotic translation initiation is tightly regulated by interactions among translation initiation factors (eIFs) that ensure accurate start codon selection. The translation regulator, eIF5 mimic protein 1 (5MP1) contributes to this process by competing with eIF5 for binding to eIF2, thereby increasing the stringency of translation initiation. Despite its important regulatory role and emerging involvement in tumorigenesis, structural information on human 5MP1 remains limited. Here, we report the near-complete backbone and partial side-chain NMR resonance assignments of the C-terminal domain of human 5MP1 (residues 250-419), carrying a W404E substitution that disrupts dimerization. The WT protein forms a dimer at NMR concentrations, which increases the effective size of the protein and also causes disappearance of peaks corresponding to aminoacids at the dimer interface due to conformational exchange. Backbone resonance assignments were completed for 96.4% of the non-proline residues. Secondary structure was analyzed using Chemical Shift Index (CSI) and compared with the AlphaFold structural model. Regions of disagreement between the experimental and computational secondary structure assignments were further examined using 15N-NOESY-HSQC spectra, allowing experimental validation of local structural features. While the AlphaFold model accurately reproduces the overall fold of the 5MP1 C-terminal domain, several localized discrepancies were identified, particularly near the N- and C-terminal regions of the domain, where experimental NMR data support alternative secondary structure assignments. These resonance assignments and experimentally validated structural features provide a foundation for future investigations of the molecular interactions, dynamics, and functions of 5MP1 in translation initiation.
Daniel, J.; Vitoriano De Queiroz Lira, L.; Zea, D. J.
Show abstract
Proteins are dynamic molecules capable of adopting multiple conformations. However, AlphaFold2 predominantly generates models around a single conformation, usually representing a ligand-bound state. To address this limitation, we developed AlphaConformers, a structure-guided pipeline that steers AlphaFold2 toward alternative conformations. It is based on the idea that protein structure databases can capture the structural space accessible to members of a protein family. Given a target protein, AlphaConformers retrieves structures from structurally similar proteins. These structures are organized into structure-based alignments and template sets, which are supplied to AlphaFold2 as conformational hypotheses. The resulting models are clustered and filtered, facilitating their analysis. Evaluated on a curated benchmark of 88 proteins with known ligand-bound and unbound conformations, AlphaConformers expanded AlphaFold2 conformational sampling and recovered alternative states missed by AlphaFold2 and other state-of-the-art methods. AlphaConformers ranked first for modelling subtle conformational changes commonly observed between ligand-bound and unbound states. These results show that structural information from protein databases can be leveraged to steer AlphaFold2 toward alternative conformations.
Zhang, Z.; Ibtehaz, N.; Kagaya, Y.; Xu, Z.; Punuru, P.; Kihara, D.
Show abstract
Recent advances in protein structure prediction, exemplified by AlphaFold, have largely addressed the determination of static structures, one aspect of the protein folding problem. However, predicting folding pathways, by which proteins reach their native states, remains a significant challenge. Here, we present PathFold, a deep learning framework that predicts protein folding pathways directly from sequence information. PathFold leverages an AlphaFold-based module to extract structural information from the sequence and generates a progressive folding trajectory from an extended conformation using a diffusion model. By modeling the full trajectory, it enables prediction of folding intermediates and transition pathways, analogous to those observed in steered molecular dynamics (SMD) simulations. The predicted pathways reveal well-defined intermediates and sequential folding events, and show agreement with experimental folding data, including measured {Phi}-values.
Buton, N.; Piochi, L. F.; Khakzad, H.
Show abstract
Intrinsically disordered proteins and regions (IDPs/IDRs) mediate diverse cellular functions through binding segments whose functional properties are encoded in dynamic conformational ensembles rather than a single static state. Existing predictors of linear interacting peptides (LIPs) and molecular recognition features (MoRFs) rely primarily on sequence-derived features, leaving ensemble-level biophysical properties largely unexplored. Here, we introduce BindCORE, an ensemble-aware deep learning framework that integrates global, local, and pairwise biophysical descriptors to predict interaction sites within IDRs. These features are processed through a multi-scale architecture that enables information exchange between sequence- and ensemble-based global, local, and pairwise information. Across established LIP and MoRF benchmarks, BindCORE consistently improves performance over sequence-based baselines, demonstrating the predictive signals of ensemble-derived properties beyond sequence-based representations alone. Feature-attribution analyses reveal that pairwise descriptors are the dominant contributors to prediction, while solvent accessibility, backbone dihedral entropy, and global geometric properties provide complementary information. Feature-importance rankings vary substantially across ensemble flavours, indicating that different conformational generators encode distinct biophysical signatures of interaction-site propensity. Together, our results show that conformational ensembles contain interpretable determinants of LIP and MoRF binding residues and establish BindCORE as a general framework for incorporating biophysical information into the prediction of functional regions in intrinsically disordered proteins. BindCORE is freely available as a ready-to-use Google Colab notebook at https://gitlab.inria.fr/delta/bindcore.
Torres, M. D. T.; Cao, H.; de la Fuente-Nunez, C.
Show abstract
Many peptides often do not have a single dominant structure. Instead, many remain disordered in water and fold when they encounter membranes or other chemical environments, a property that underlies diverse biological functions but is difficult to predict. Here we introduce ApexFold, a machine-learning frame-work that predicts how peptide secondary structure change across environments. ApexFold uses peptide sequence and features together with physicochemical descriptors of the surrounding medium to estimate the fractions of helical, {beta}-like and disordered structure expected in each condition. Trained on circular-dichroism measurements from 1,187 peptides assayed in water, co-solvents and membrane-mimicking micelles, ApexFold predicted solvent-induced structural shifts in independent peptide panels and outperformed static structure predictors that return a single conformation. These results show that peptide structural plasticity can be learned from sequence and environment, providing a way to prioritize peptides and experimental conditions before synthesis and structural characterization.